Papers with multimodal machine translation model

2 papers
Domain Adaptation of Image Encoder for Multimodal Manga Translation (2026.eacl-srw)

Copied to clipboard

Challenge: Existing machine translation systems lack sufficient manga comprehension capabilities when utilizing image information.
Approach: They propose a domain-adapted image encoder training method for manga . the method trains encoders to acquire visual features that consider the structural and sequential characteristics of the manga based on a Japanese-English translation task.
Outcome: The proposed method improves translation evaluation metrics in Japanese-English translation task compared to the conventional method .
A Visual Attention Grounding Neural Model for Multimodal Machine Translation (D18-1)

Copied to clipboard

Challenge: Existing approaches to multimodal machine translation do not integrate visual information into the translation process.
Approach: They propose a multimodal machine translation model that utilizes parallel visual and textual information.
Outcome: The proposed model outperforms existing methods on the Multi30K and Ambiguous COCO datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations